Papers with misinformation detection
Detection and Resolution of Rumors and Misinformation with NLP (2020.coling-tutorials)
Copied to clipboard
| Challenge: | Detecting false and misleading claims on the web is a sub-field of NLP . this half-day tutorial presents the theory behind each of these steps and the state-of-the-art solutions. |
| Approach: | This half-day tutorial presents the theory behind false and misleading claims detection . it covers the steps involved in identifying check-worthy claims, tracking claims and rumors, rumor collection and annotation, grounding claims against knowledge bases, and using stance to verify claims. |
| Outcome: | This half-day tutorial presents the theory behind each of these steps and the state-of-the-art solutions. |
Fairness-Aware Online Positive-Unlabeled Learning (2024.emnlp-industry)
Copied to clipboard
| Challenge: | Positive-unlabeled (PU) learning is a new approach to improve text classification by analyzing the impact of the online setting on fairness. |
| Approach: | They propose to extend Positive-Unlabeled (PU) learning to online learning by analyzing the impact of the online setting on fairness. |
| Outcome: | The proposed approach improves fairness in PU learning in both offline and online settings by using only labeled positive and unlabeled samples. |
ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing efforts in misinformation detection focus on written text, leaving a significant gap in addressing the complexity of spoken text in video transcripts. |
| Approach: | They propose to annotate video transcripts in three languages and six topics using a custom annotation tool. |
| Outcome: | The proposed tool shows strong cross-validation performance but challenges for generalization to unseen topics. |
MMM: An Emotion and Novelty-aware Approach for Multilingual Multimodal Misinformation Detection (2022.findings-aacl)
Copied to clipboard
| Challenge: | Increasing presence of multimedia content on the web promotes misinformation . detecting this category of misleading information is almost impossible without prior knowledge . |
| Approach: | They propose a novel multilingual multimodal misinformation dataset that includes background knowledge of misleading articles. |
| Outcome: | The proposed model outperforms the state-of-the-art on misinformation detection task. |
o-MEGA: Optimized Methods for Explanation Generation and Analysis (2025.emnlp-demos)
Copied to clipboard
| Challenge: | a growing number of transformer-based language models have created challenges for model transparency and trustworthiness. |
| Approach: | They propose a tool to automatically identify the most effective explainable AI methods . they evaluate o-mega on a post-claim matching pipeline using a curated dataset . |
| Outcome: | The proposed tool shows that the most effective explainable AI methods can be implemented in semantic matching tasks. |
On the Risk of Misinformation Pollution with Large Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study demonstrates that large language models can be misused for generating credible-sounding misinformation . however, the ability to produce credible text raises concerns regarding their potential misuse . |
| Approach: | They propose three defense strategies to mitigate misinformation generated by Large Language Models . they propose a threat model and simulate potential misuse scenarios . |
| Outcome: | The proposed defense strategies have shown promising results, albeit with costs. |
No Innocence in Styling: Discovery of Privacy Protection Capabilities and Security Risks in Consumer Generative AI Writing Assistants (2026.acl-industry)
Copied to clipboard
Mohd. Farhan Israk Soumik, Syed Mhamudul Hasan, Wanniarachchi Kankanamge Malithi Mithsara, Ahmed Imteaj, Abdur R. Shahid
| Challenge: | a recent study examines the dual-use nature of platform-level text stylization. |
| Approach: | They examine the dual-use nature of platform-level text stylization by examining their implications for privacy and platform safety. |
| Outcome: | The proposed model reduces emotion inference accuracy, lowers profiling risk, and increases error rates in misinformation detection. |
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models are limited by challenges in factuality and hallucinations to be directly employed off-the-shelf for judging the veracity of news articles. |
| Approach: | They propose to integrate large language models into the news pipeline by generating news reactions and generating proxy tasks. |
| Outcome: | The proposed model outperforms state-of-the-art baselines by 16.8% in macro f1-score on seven datasets with three LLMs. |
Efficient Annotator Reliability Assessment and Sample Weighting for Knowledge-Based Misinformation Detection on Social Media (2025.findings-naacl)
Copied to clipboard
Owen Cook, Charlie Grimshaw, Ben Peng Wu, Sophie Dillon, Jack Hicks, Luke Jones, Thomas Smith, Matyas Szert, Xingyi Song
| Challenge: | Misinformation spreads rapidly on social media, confusing the truth and targeting potentially vulnerable people. |
| Approach: | They propose to use inter- and intra-annotator agreement to understand the reliability of each annotator and influence the training of large language models based on annotators reliability. |
| Outcome: | The proposed framework utilises inter- and intra-annotator agreement to understand the reliability of each annotator and influence the training of large language models based on annotators reliability. |
MetaAdapt: Domain Adaptive Few-Shot Misinformation Detection via Meta Learning (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for misinformation detection are limited by data scarcity . existing methods fail to detect early-stage misinformation on emerging topics . |
| Approach: | They propose a meta learning based approach for domain adaptive few-shot misinformation detection that leverages limited target examples to provide feedback and guide the knowledge transfer from the source to the target domain. |
| Outcome: | The proposed method improves performance on real-world datasets with reduced parameters. |
Multimodal Fact-Checking with Vision Language Models: A Probing Classifier based Solution with Embedding Strategies (2025.coling-main)
Copied to clipboard
| Challenge: | Existing fact-checking systems that use text and image information are susceptible to fake news spread by social media platforms. |
| Approach: | They propose a neural probing classifier based on multimodality and embeddings from text and image encoders to represent multimodal content for fact-checking. |
| Outcome: | The proposed classifier outperforms KNN and SVM baselines in leveraging extracted embeddings, highlighting its effectiveness for multimodal fact-checking. |
MM-SOC: Benchmarking Multimodal Large Language Models in Social Media Platforms (2024.findings-acl)
Copied to clipboard
| Challenge: | Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in online spaces. |
| Approach: | They propose a benchmark to evaluate MLLMs' understanding of multimodal social media content and a large-scale YouTube tagging dataset to evaluate their performance. |
| Outcome: | The proposed model performs better in a zero-shot setting, suggesting potential improvements. |
The Pragmatics behind Politics: Modelling Metaphor, Framing and Emotion in Political Discourse (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing computational models of political discourse do not incorporate metaphor and emotion in their functions. |
| Approach: | They propose to combine metaphor, emotion and political rhetoric to model political discourse . they show that they advance in three tasks: predicting political perspective of news articles, party affiliation of politicians and framing of policy issues. |
| Outcome: | The proposed models improve political discourse prediction, party affiliation and framing of policy issues. |
The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection (2025.emnlp-main)
Copied to clipboard
Arghodeep Nandi, Megha Sundriyal, Euna Mehnaz Khan, Jikai Sun, Emily K. Vraga, Jaideep Srivastava, Tanmoy Chakraborty
| Challenge: | a survey examines the interplay between factual accuracy and cognitive biases . misinformation is more than just the existence of incorrect information, it also entails complex relationships between the information and the entities that consume it. |
| Approach: | They examine the interplay between traditional fact-checking and psychological concepts such as cognitive biases, social dynamics, and emotional responses. |
| Outcome: | The findings highlight limitations of current methods and identify opportunities for improvement . they also outline future research directions to create more robust frameworks . |
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection (2025.findings-acl)
Copied to clipboard
| Challenge: | Multimodal large language models have demonstrated impressive capabilities in visual reasoning and text generation. |
| Approach: | They propose a multimodal large language model that captures deeper relationships between images and text . they propose CMIE, which uses a Coexistence Relationship Generation strategy and an AS mechanism to detect misinformation. |
| Outcome: | The proposed framework outperforms existing methods in detecting out-of-context misinformation. |
LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection (2026.acl-long)
Copied to clipboard
| Challenge: | Current evaluation frameworks are static and vulnerable to benchmark data contamination . current models are ineffective at assessing reasoning under temporal uncertainty . |
| Approach: | They propose a live-based benchmark that simulates the real-world "fog of war" they propose evaluating models on their ability to reason with evolving, incomplete information . |
| Outcome: | The proposed model outperforms proprietary state-of-the-art models in classification and evidence mode . it also provides a component to monitor BDC explicitly . |
A Zero-Shot Claim Detection Framework Using Question Answering (2022.coling-1)
Copied to clipboard
| Challenge: | Existing claims detection frameworks are portability to emerging events and low-resource training data settings. |
| Approach: | They propose a claim detection framework that leverages zero-shot Question Answering to solve sub-tasks such as topic filtering, claim object detection, and claimer detection. |
| Outcome: | The proposed framework outperforms baselines on the NewsClaims benchmark. |
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models (2023.acl-long)
Copied to clipboard
| Challenge: | Hundreds of studies have highlighted ethical issues in NLP models . |
| Approach: | They propose to measure media biases in LMs trained on diverse data sources . they focus on hate speech and misinformation detection . |
| Outcome: | The proposed methods quantify the fairness of downstream NLP models trained on politically biased LMs. |
SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness (2024.emnlp-main)
Copied to clipboard
Tanmay Parekh, Jeffrey Kwan, Jiarui Yu, Sparsh Johri, Hyosang Ahn, Sreya Muppalla, Kai-Wei Chang, Wei Wang, Nanyun Peng
| Challenge: | Prior studies focused on English posts to provide early warnings for epidemic prediction, but these work focused on non-English posts. |
| Approach: | They propose a multilingual event extraction framework for extracting epidemic event information for any disease and language using 5.1K tweets in four languages. |
| Outcome: | The proposed framework can provide epidemic warnings for COVID-19 in its earliest stages in Dec 2019 (3 weeks before global discussions) and aggregate community epidemic discussions like symptoms and cure measures, aiding misinformation detection and public attention monitoring. |
Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Despite advances in large language models, their application to misinformation detection remains hindered by issues of logical inconsistency and superficial verification. |
| Approach: | They propose a multi-agent debate framework that reformulates misinformation detection as a structured adversarial debate based on fact-checking workflows . |
| Outcome: | The proposed framework enables iterative refinement of evidence while improving decision transparency. |
RAEmoLLM: Retrieval Augmented LLMs for Cross-Domain Misinformation Detection Using In-Context Learning Based on Emotional Information (2025.acl-long)
Copied to clipboard
| Challenge: | Current methods for cross-domain misinformation detection focus on in-domain tasks and do not incorporate significant sentiment and emotion features. |
| Approach: | They propose a retrieval augmented (RAG) LLM framework that incorporates affective information into retrieval databases. |
| Outcome: | The proposed framework improves on three misinformation benchmarks. |
Misinformation with Legal Consequences (MisLC): A New Task Towards Harnessing Societal Harm of Misinformation (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing research has focused on the veracity of information, overlooking the legal implications and consequences of misinformation. |
| Approach: | They propose a task to detect misinformation using legal issues as a measure of societal ramifications. |
| Outcome: | The proposed task leverages definitions from a wide range of legal domains covering 4 broader legal topics and 11 fine-grained legal issues, including hate speech, election laws, and privacy regulations. |
MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | a new dataset of misinformation labels is being developed to detect misinformation on social media platforms . misinformation is spread in many domains including but not limited to health, politics, and disasters . |
| Approach: | They construct a dataset of 5,284 English and 5,064 Turkish tweets with misinformation labels . they use the dataset to analyze misinformation spread and to evaluate misinformation detection . |
| Outcome: | The proposed dataset includes 5,284 English and 5,064 Turkish tweets with misinformation labels for several recent events between 2020 and 2022. |
Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback (2025.findings-emnlp)
Copied to clipboard
| Challenge: | popular training paradigms for language models often assume there is one optimal answer for every query. |
| Approach: | They propose to enhance pluralistic alignment of language models using pluralistic decoding and model steering methods. |
| Outcome: | The proposed methods improve pluralistic alignment of language models in a low-resource setting . the proposed methods decrease false positives in several high-stakes tasks . |
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Infodemics and health misinformation have significant negative impact on individuals and society . generative AI has significantly accelerated the spread and expanded the reach of health misinfo . |
| Approach: | MM-Health is a large scale multimodal misinformation dataset in the health domain . it includes human-generated multimodal information and AI-generated multiplemodal information . |
| Outcome: | MM-Health is a large scale misinformation dataset in the health domain . it includes human-generated multimodal information and AI-generated content . |
Using Persuasive Writing Strategies to Explain and Detect Health Misinformation (2024.lrec-main)
Copied to clipboard
| Challenge: | Increasing misinformation has led to a decrease in trust in news organizations and a decline in the health and medical industry. |
| Approach: | They propose a novel annotation scheme that incorporates persuasive writing tactics in textual documents to aid the automatic identification of misinformation. |
| Outcome: | The proposed scheme improves accuracy and explainability of misinformation detection models. |